Papers with statistical analyses

5 papers
EfficientOCR: An Extensible, Open-Source Package for Efficiently Digitizing World Knowledge (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing OCR engines fail to provide accurate, cost-effective and sample-efficient character recognition for public domain documents.
Approach: EffOCR is an open-source optical character recognition package that is accurate, cheap to deploy and sample efficient to customize to novel collections, languages, and character sets.
Outcome: EffOCR model trains character retrieval problem and scales to novel collections, languages, and character sets.
Does Generative AI speak Nigerian-Pidgin?: Issues about Representativeness and Bias for Multilingualism in LLMs (2025.findings-naacl)

Copied to clipboard

Challenge: Nigeria is a multilingual country with 500+ languages.
Approach: They propose to use a pidgin and a creole to analyze the pidgins of Nigeria . they also use machine translation to analyze their results .
Outcome: The results show that the two pidgins do not represent each other and are hard to teach . the results show the pidgin varieties are underrepresented in Generative AI .
Replicating and Extending “Because Their Treebanks Leak”: Graph Isomorphism, Covariants, and Parser Performance (2021.acl-short)

Copied to clipboard

Challenge: a small sample size and unreliable results suggest a correlation between parser performance and graph isomorphism is not observed in the wild.
Approach: They propose to replicate a study which found graph isomorphism is a non-trivial variable . they also bin sentences by length and find correlation between parser performance and isopathism disappears .
Outcome: The results show that the original analysis was unreliable and had methodological issues . the study also bin sentences by length and shows that the correlation between parser performance and graph isomorphism disappears when controlling for covariants.
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented.
Approach: They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods.
Outcome: The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents.
ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media (2025.emnlp-main)

Copied to clipboard

Challenge: Almost 50% of depression patients face the risk of going into relapse.
Approach: They propose to validate a social media dataset on depression relapse using cognitive theories of depression.
Outcome: The first clinically validated social media dataset focused on depression relapse comprises 204 Reddit users annotated by mental health professionals.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations